Papers by Keng Ji Chow
TraVLR: Now You See It, Now You Don’t! A Bimodal Dataset for Evaluating Visio-Linguistic Reasoning (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing visio-linguistic (V+L) models do not represent visual and linguistic concepts in a unified space. |
| Approach: | They propose to use cross-modal transfer to evaluate the extent to which visio-linguistic (V+L) representations are represented in a unified space. |
| Outcome: | The proposed evaluation settings include cross-modal transfer and a global accuracy score on the entire dataset making the specific sources of success and failure difficult to diagnose. |